Detecting genomic indel variants with exact breakpoints in single- and paired-end sequencing data using SplazerS

نویسندگان

  • Anne-Katrin Emde
  • Marcel H. Schulz
  • David Weese
  • Ruping Sun
  • Martin Vingron
  • Vera M. Kalscheuer
  • Stefan A. Haas
  • Knut Reinert
چکیده

MOTIVATION The reliable detection of genomic variation in resequencing data is still a major challenge, especially for variants larger than a few base pairs. Sequencing reads crossing boundaries of structural variation carry the potential for their identification, but are difficult to map. RESULTS Here we present a method for 'split' read mapping, where prefix and suffix match of a read may be interrupted by a longer gap in the read-to-reference alignment. We use this method to accurately detect medium-sized insertions and long deletions with precise breakpoints in genomic resequencing data. Compared with alternative split mapping methods, SplazerS significantly improves sensitivity for detecting large indel events, especially in variant-rich regions. Our method is robust in the presence of sequencing errors as well as alignment errors due to genomic mutations/divergence, and can be used on reads of variable lengths. Our analysis shows that SplazerS is a versatile tool applicable to unanchored or single-end as well as anchored paired-end reads. In addition, application of SplazerS to targeted resequencing data led to the interesting discovery of a complete, possibly functional gene retrocopy variant. AVAILABILITY SplazerS is available from http://www.seqan.de/projects/ splazers. SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

I-37: Establishing High Resolution Genomic Profiles of Single Cells Using Microarray and Next-Generation Sequencing Technologies

The nature and pace of genome mutation is largely unknown. Standard methods to investigate DNA-mutation rely on arraying or sequencing DNA from a population of cells, hence the genetic composition of individual cells is lost and de novo mutation in cell(s) is concealed within the bulk signal. We developed methods based on (SNP-) arraying and next-generation sequencing of single-cell whole-genom...

متن کامل

Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads

MOTIVATION There is a strong demand in the genomic community to develop effective algorithms to reliably identify genomic variants. Indel detection using next-gen data is difficult and identification of long structural variations is extremely challenging. RESULTS We present Pindel, a pattern growth approach, to detect breakpoints of large deletions and medium-sized insertions from paired-end ...

متن کامل

PRISM: Pair-read informed split-read mapping for base-pair level detection of insertion, deletion and structural variants

MOTIVATION The development of high-throughput sequencing technologies has enabled novel methods for detecting structural variants (SVs). Current methods are typically based on depth of coverage or pair-end mapping clusters. However, most of these only report an approximate location for each SV, rather than exact breakpoints. RESULTS We have developed pair-read informed split mapping (PRISM), ...

متن کامل

Seeksv: an accurate tool for somatic structural variation and virus integration detection

MOTIVATION Many forms of variations exist in the human genome including single nucleotide polymorphism, small insert/deletion (DEL) (indel) and structural variation (SV). Somatically acquired SV may regulate the expression of tumor-related genes and result in cell proliferation and uncontrolled growth, eventually inducing tumor formation. Virus integration with host genome sequence is a type of...

متن کامل

A Method to Detect Large Deletions at Base Pair Level using Next-Generation Sequencing Data

Genomic structural variants (SV) play a significant role in the onset and progression of cancer. Genomic deletions can create oncogenic fusion genes or cause the loss of tumor suppressing gene function which can lead to tumorigenesis by downregulating these genes. Detecting these variants has clinical importance in the treatment of diseases. Furthermore, it is also clinically important to detec...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:
  • Bioinformatics

دوره 28 5  شماره 

صفحات  -

تاریخ انتشار 2012